Papers with predictive accuracy

44 papers
Augur: Modeling Covariate Causal Associations in Time Series via Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have emerged as a promising avenue for time series forecasting . existing approaches face limitations such as marginalized role in model architectures and lack of interpretability.
Approach: They propose a framework that exploits LLM causal reasoning to discover and use directed causal associations among covariates.
Outcome: The proposed model improves predictive accuracy while yielding transparent, traceable reasoning about variable interactions.
Literature-Augmented Clinical Outcome Prediction (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to clinical outcome prediction use only clinical notes and general biomedical literature.
Approach: They propose to retrieve patient-specific medical literature and incorporate it into predictive models by combining clinical notes with language models.
Outcome: The proposed approach boosts predictive performance on three important clinical tasks in comparison to strong LM baselines, increasing F1 by up to 5 points and precision@Top-K by a large margin of over 25%.
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models (2026.eacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque.
Approach: They propose an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an explanation-driven process.
Outcome: The proposed framework produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines.
Disentangling Representations of Text by Masking Transformers (2021.emnlp-main)

Copied to clipboard

Challenge: Large pretrained models such as BERT encode a range of features into monolithic vectors, providing strong predictive accuracy across downstream tasks.
Approach: They explore whether it is possible to learn disentangled representations by identifying existing subnetworks within pretrained models that encode distinct, complementary aspects.
Outcome: The proposed method disentangles sentiment from genre in movie reviews, toxicity from dialect in Tweets, and syntax from semantics.
Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys (2024.findings-eacl)

Copied to clipboard

Challenge: integrating cultural dimensions with dialogue encoding features can enhance the predictive accuracy and quality of dialogue agents.
Approach: They propose to incorporate cultural dimensions into dialogue encoding features to enhance the predictive accuracy of dialogue agents.
Outcome: The proposed model improves the accuracy and quality of dialogue predictions by incorporating cultural dimensions with dialogue encoding features.
Diverse Few-Shot Text Classification with Multiple Metrics (N18-1)

Copied to clipboard

Challenge: Existing methods for few-shot learning are insufficient to capture task variations in natural language domains.
Approach: They propose an adaptive metric learning approach that automatically determines the best weighted combination from a set of metrics obtained from meta-training tasks for a newly seen few-shot task.
Outcome: The proposed method performs favorably against state-of-the-art few shot learning algorithms on real-world sentiment analysis and dialog intent classification datasets.
RBPtool: A Deep Language Model Framework for Multi-Resolution RBP-RNA Binding Prediction and RNA Molecule Design (2025.emnlp-main)

Copied to clipboard

Challenge: RNA-binding proteins play key roles in post-transcriptional gene regulation . existing methods focus on shallow sequence features or coarse structural representations . large language models allow for precise modeling and biologically informed de novo RNA design .
Approach: They extend RPI15223 into a multi-resolution, structure-level RBP-RNA dataset and introduce RBPtool, a framework that fuses sequence and structural information.
Outcome: The proposed framework achieves state-of-the-art performance on public benchmarks and the RPI15223 dataset while supporting fine-grained level predictions.
SparseFit: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations (2024.acl-long)

Copied to clipboard

Challenge: Models that generate natural language explanations (NLEs) for their predictions often require large datasets of human-written NLEs at training time, which can be expensive and time-consuming to collect.
Approach: They propose a sparse few-shot fine-tuning strategy that leverages discrete prompts to jointly generate predictions and NLEs.
Outcome: The proposed approach compares sparse few-shot fine-tuning with existing parametric fine- tuning techniques on three sizes of the T5 language model and four datasets and produces competitive results for both task performance and NLE quality.
Surge: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors (2025.emnlp-main)

Copied to clipboard

Challenge: Neural surrogate models are powerful tools in data mining, but are underexplored . large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks .
Approach: They propose a benchmarking framework to examine the feasibility of large language models . they examine scaling laws, data efficiency, and predictive accuracy of 21 open-source and proprietary LLMs .
Outcome: The proposed benchmark examines 21 open-source and proprietary LLMs . it also examines scaling laws, data efficiency, and predictive accuracy .
I Beg to Differ: A study of constructive disagreement in online conversations (2021.eacl-main)

Copied to clipboard

Challenge: Disagreements are pervasive in human communication.
Approach: They construct a corpus of Wikipedia Talk page conversations that contain content disputes and define the task of predicting whether disagreements will be escalated to mediation by a moderator.
Outcome: The proposed model outperforms feature-based models in predicting whether disagreements will escalate to mediation by a moderator.
Interpretable Graph-Language Modeling for Detecting Youth Illicit Drug Use (2026.findings-eacl)

Copied to clipboard

Challenge: Illicit drug use among teens and young adults remains a public health concern . existing models ignore latent and interconnected structures among survey variables .
Approach: They propose a joint graph-language modeling framework to detect illicit drug use among TYAs . they use large-scale surveys such as the Youth Risk Behavior Survey and the National Survey on Drug Use and Health to analyze data .
Outcome: The proposed framework outperforms baseline models on YRBS and NSDUH datasets in predictive accuracy.
Unleashing the Potential of Large Language Models through Spectral Modulation (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, garnering significant attention from both academia and industry.
Approach: They propose to conduct spectral modulation in the parameter space of LLMs to integrate with various models in a plug-and-play manner.
Outcome: The proposed approach improves performance by 10.12% with spectral modulation.
Logical Reasoning with Span-Level Predictions for Interpretable and Robust NLI Models (2022.emnlp-main)

Copied to clipboard

Challenge: Current models learn from annotation artefacts and dataset biases, but it is unclear to what extent they are learning the task of NLI.
Approach: They propose a logical reasoning framework that allows models to learn from annotation artefacts and dataset biases.
Outcome: The proposed model outperforms humans on in-distribution test sets without using span labels . the model is more robust in a reduced data setting, and out-of-disturbance performance is improved .
Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts (2024.findings-acl)

Copied to clipboard

Challenge: Prediction with Explanation is the largest expert-annotated dataset for legal judgment prediction and explanation in the Indian context .
Approach: They propose to use an annotated legal judgment prediction corpus to improve models' accuracy . they employ transformer-based models tailored for both general and Indian legal contexts .
Outcome: The proposed system improves the accuracy and explanatory depth of models for legal judgments.
Evaluating the Utility of Hand-crafted Features in Sequence Labelling (D18-1)

Copied to clipboard

Challenge: Conventional wisdom is that hand-crafted features are redundant for deep learning models . authors propose a method for using handcrafted features in a hybrid learning approach .
Approach: They propose a method for exploiting handcrafted features as part of a hybrid learning approach.
Outcome: The proposed method outperforms baseline models on a named entity recognition task and reduces training requirements to 60% while maintaining the same predictive accuracy.
Exploiting Careful Design of SVM Solution for Aspect-term Sentiment Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Aspect-term sentiment analysis (ATSA) identifies fine-grained sentiments towards specific aspects of text.
Approach: They propose a pipeline to predict fine-grained sentiments for specific aspects of text . it decomposes the learning problem into multiple view subproblems and dynamically selects and constructs features with reinforcement learning.
Outcome: The proposed pipeline surpasses SVM-based methods in predictive accuracy while maintaining a faster inference speed and significantly reducing the number of model parameters.
An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods rely on parametric models that store knowledge in parameters or retrieval-augmented models that have access to external knowledge sources.
Approach: They propose a parametric parametric model that stores knowledge in its parameters or a retrieval-augmented model that has access to external knowledge sources.
Outcome: The proposed method runs substantially faster across the board and produces more accurate results on WoW and ELI5.
Mitigating Data Poisoning in Text Classification with Differential Privacy (2021.findings-emnlp)

Copied to clipboard

Challenge: Data poisoning attacks can plant a backdoor in a model by injecting poisoned examples into training data, causing the model to misclassify test instances which include a specific pattern.
Approach: They propose a generic defence mechanism that makes training robust to poisoning attacks by smoothing the gradient from each training example.
Outcome: The proposed method is highly effective in mitigating, or even eliminating, poisoning attacks on text classification, with only a small cost in predictive accuracy.
PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements (2025.emnlp-main)

Copied to clipboard

Challenge: Contract review is a complex and time-intensive task that typically requires legal expertise.
Approach: a new open-source contract review framework is designed to handle complexities of contract analysis . PAKTON is a retrieval-augmented generation framework with plug-and-play capabilities .
Outcome: The open-source framework outperforms models in predictive accuracy, retrieval performance, explainability, completeness, and grounded justifications.
Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control (D19-1)

Copied to clipboard

Challenge: Selective rationalization is a common mechanism to ensure that predictive models reveal how they use any available features.
Approach: They propose a co-operative method which uses introspection to explicitly predict and incorporate the outcome into the selection process.
Outcome: The proposed model maintains high predictive accuracy and leads to comprehensive rationales.
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies highlight the effectiveness of using in-context learning (ICL) to steer large language models in processing tabular data.
Approach: They propose a method that uses clustering and evolutionary strategies to curate a representative sample set from training data.
Outcome: The proposed method significantly improves fairness across various metrics, showing its efficacy in real-world scenarios.
Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability.
Approach: They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality.
Outcome: The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks.
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models (2025.coling-main)

Copied to clipboard

Challenge: a study examines the behavior of routers in Mixture of Experts (MoE) models . experts with similar linguistic traits are often routed to the same expert regardless of context .
Approach: They investigate how tokens are routed based on their linguistic features . they aim to explore whether experts specialize in processing tokens with similar linguistic traits .
Outcome: The proposed model-integrated routers are based on Mixture of Experts (MoE) models . the results show that expert specialization is high for POS categories .
Learning to Deceive with Attention-Based Explanations (2020.acl-main)

Copied to clipboard

Challenge: Attention mechanisms are ubiquitous components in neural network architectures and are often claimed to confer interpretability.
Approach: They propose a method for training models to produce deceptive attention masks by combining weights assigned to designated impermissible tokens with a weighted sum.
Outcome: The proposed method reduces the weight assigned to designated impermissible tokens while still using them across multiple models and tasks.
Foiling Training-Time Attacks on Neural Machine Translation Systems (2022.findings-emnlp)

Copied to clipboard

Challenge: Neural machine translation systems are vulnerable to backdoor attacks . successful backdoors can cause slander, hate speech, phishing, etc. attacks can target very short trigger phrases, which can be challenging to detect even when included verbatim in poisoned instances.
Approach: They propose a method that exploits asymmetry between source and target sentences to detect outlier tokens.
Outcome: The proposed method reduces the success of attacks by up to 89.0% while not affecting predictive accuracy.
Learning Constraints for Structured Prediction Using Rectifier Networks (2020.acl-main)

Copied to clipboard

Challenge: Various natural language processing tasks require domain expertise to design good constraints.
Approach: They propose a framework for learning constraints in a network of linear inequalities over the output variables.
Outcome: The proposed framework can be used to learn constraints from data on natural language processing tasks.
Efficient Integration of External Knowledge to LLM-based World Models via Retrieval-Augmented Generation and Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing attempts to enhance LLM-based world models through prompting or fine-tuning approaches are either requiring human knowledge or computationally extensive.
Approach: They propose a framework that leverages retrieval-augmented generation to integrate external knowledge to LLM-based world models.
Outcome: The proposed framework outperforms baseline models and exhibits strong generalizability.
Cognitive Alpha Mining via LLM-Driven Code-Based Evolution (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to finding effective predictive signals from financial data are limited by their complexity and low signal-to-noise ratio.
Approach: They propose a framework that combines code-level alpha representation with LLM-driven reasoning and evolutionary search.
Outcome: The proposed framework combines code-level alpha representation with LLM-driven reasoning and evolutionary search.
Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing models struggle to balance predictive accuracy with human-understandable rationales.
Approach: They propose to enhance LLMs by leveraging rationale distillation and domain knowledge injection for trustworthy multimodal rationale generation.
Outcome: Experiments on real-world medical datasets show that ClinRaGen achieves state-of-the-art performance in disease diagnosis and rationale generation.
Are representations built from the ground up? An empirical examination of local composition in language models (2022.emnlp-main)

Copied to clipboard

Challenge: Compositionality is a hallmark of human language, but many phrases are non-compositional . a study by a team of researchers shows that LMs may not be able to distinguish between compositional and non-composable phrases.
Approach: They propose to predict LM-internal representations of longer phrases given their constituents . they find that the representation of a parent phrase can be predicted with some accuracy .
Outcome: The proposed model can predict a parent phrase with some accuracy given its children's transformations, but this is not the case.
Semantic and Sentiment Dual-Enhanced Generative Model for Script Event Prediction (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to model event associations struggle with semantic ambiguity and embedding bias.
Approach: They propose a Semantic and Sentiment Dual-enhanced Generative Model to address these issues . it leverages two types of script event information to enhance the generative model .
Outcome: The proposed model captures both global and local sentiments of events through its sentiment awareness mechanism.
STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing models focus on predictive accuracy over reasoning, a gap exists . time series data are ubiquitous in real-world systems and exhibit complex spatio-temporal structures.
Approach: They propose a time series reasoning model that integrates time series, graph structure, and text for explicit reasoning.
Outcome: The proposed model achieves average accuracy gains between 17% and 135% at 0.004x the cost of proprietary models and generalizes robustly to real-world data.
Prediction-Augmented Generation for Automatic Diagnosis Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) adopt autoregressive architecture, predicting the next word token based on the preceding context.
Approach: They propose a method that integrates task-specific predictive models as external tools to improve model generation quality and accuracy.
Outcome: The proposed method improves the generation quality and predictive accuracy of large language models in inference-driven tasks.
Measuring Consistency in Text-based Financial Forecasting Models (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have allowed financial forecasting to gain significant accuracy and reliability.
Approach: They propose a tool that assesses logical consistency in financial text and compares it with other models to assess their performance.
Outcome: The proposed evaluation tool assesses logical consistency in financial text.
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling (2026.acl-long)

Copied to clipboard

Challenge: Existing research on inference scaling focuses on unstructured output generation tasks, such as mathematical problems.
Approach: They propose an inference-scaling framework that combines fine-grained beam search with ToolPRM, a process reward model scoring each intra-call decision.
Outcome: The proposed framework outperforms outcome and coarse-grained reward models in predictive accuracy and yields consistent test-time gains on multiple function-calling benchmarks.
QuantAgents: Towards Multi-agent Financial System via Simulated Trading (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing LLM-based agent models exhibit significant deviations from real-world fund companies.
Approach: They propose a multi-agent financial system that incorporates simulated trading . they propose simulated trades are evaluated without assuming actual risks .
Outcome: The proposed system evaluates various investment strategies without assuming actual risks without involving real-world investors.
CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Knowledge Tracing (KT) aims to model a student’s learning state over time and predict their future performance.
Approach: They propose a framework that harnesses Large Language Models to enhance both prediction accuracy and explainability by a synergistic optimization loop.
Outcome: The proposed framework improves both prediction accuracy and explainability by using a synergistic optimization loop.
COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto Frontier (2025.emnlp-main)

Copied to clipboard

Challenge: Prior exemplar selection methods focus on maximizing predictive accuracy, neglecting model calibration.
Approach: They propose to use a Bayesian optimization algorithm to optimize for predictive accuracy and calibration.
Outcome: The proposed algorithm beats or matches baselines on multiple tasks from un-saturated MMLU-pro benchmarks while requiring minimal number of API calls.
PsyPath: Psychologically-guided Self-Exploration for Personality Detection (2026.findings-acl)

Copied to clipboard

Challenge: Personality detection aims to label traits via identifying linguistic cues from written text.
Approach: They propose a framework that allows large language models to generate and answer psychologically meaningful questions and a hybrid scoring mechanism to evaluate the generated nodes in the reasoning paths.
Outcome: The proposed framework outperforms baselines on two benchmark datasets and significantly improves performance and interpretability in downstream tasks.
Deconstruct, Diagnose, and Deliberate: A Protocol-Adaptive Role-Specific Multi-Agent Framework for Fake News Detection (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for fake news detection rely on monolithic verification methods . Existing approaches often yield ambiguous verdicts due to superficial processing .
Approach: They propose a protocol-adaptive role-specific multi-agent framework that decomposes verification into factual, logical, and contextual dimensions.
Outcome: The proposed framework outperforms baseline methods in both predictive accuracy and explanatory quality.
Efficient Hyperparameter Optimization for LLM Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing hyperparameter optimization methods are inefficient in reinforcement learning due to model scale and resource-intensive training cycles.
Approach: They propose a hyperparameter optimization method that adapts both model size and training budget as fidelity.
Outcome: The proposed method significantly improves the computational efficiency of each trial (up to 14.9) over existing HPO methods.
The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit (2025.acl-long)

Copied to clipboard

Challenge: Existing frameworks for Large Language Models (LLMs) for Click-Through Rate prediction require a careful balance between computational efficiency and predictive accuracy.
Approach: They propose a framework that integrates Retrieval-Augmented Generation with a novel multi-head early exit architecture to address both challenges.
Outcome: The proposed framework reduces retrieval time while maintaining high model performance.
Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy (2026.findings-acl)

Copied to clipboard

Challenge: a limited amount of annotated data has slowed progress in machine learning for low-resource languages . a sentiment label records an annotator's final decision, but it is not a valid record of the annotation's interpretation.
Approach: They propose a large-scale Telugu sentiment classification dataset annotated with sentiment labels and human-selected rationales from multiple native speakers.
Outcome: The proposed model improves classification performance, explanation quality, and social bias by incorporating human rationales.
SUE: Sparsity-based Uncertainty Estimation via Sparse Dictionary Learning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to estimate uncertainty use predictive confidence, structural characteristics of representation space, or stochastic variation in model outputs.
Approach: They propose a new uncertainty estimation framework based on sparse dictionary learning by identifying dictionary atoms associated with misclassified samples.
Outcome: The proposed framework outperforms or matches existing methods on several NLU benchmarks and sentiment analysis benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations